Papers with processing pipeline
lingvis.io - A Linguistic Visual Analytics Framework (P19-3)
Copied to clipboard
Mennatallah El-Assady, Wolfgang Jentner, Fabian Sperrle, Rita Sevastjanova, Annette Hautli-Janisz, Miriam Butt, Daniel Keim
| Challenge: | Using a modular framework, linguistic visual analytics applications can be rapidly prototypized using a web-based framework. |
| Approach: | They propose a modular framework for rapid prototyping of linguistic, web-based, visual analytics applications. |
| Outcome: | The proposed framework supports rapid prototyping of linguistic, web-based, visual analytics applications. |
Revisiting Character-Based Neural Machine Translation with Capacity and Compression (D18-1)
Copied to clipboard
| Challenge: | Translating characters instead of words or word-fragments can simplify the processing pipeline but results in longer sequences . |
| Approach: | They propose to use sequence-to-sequence architectures of sufficient depth to solve the problem . they also evaluate the performance versus computation time tradeoffs they offer . |
| Outcome: | The proposed models outperform models operating over word fragments in character-level NMT, the authors show . they also show that the proposed models do not match the performance of their deep character baseline model . |
Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics (2020.lrec-1)
Copied to clipboard
| Challenge: | Swiss-AL is a multilingual web corpus for Applied Linguistics that supports data-based and data-driven research on societal and political discourses in Switzerland. |
| Approach: | They propose a multilingual Swiss web corpus for Applied Linguistics that supports data-based research on societal and political discourses in Switzerland. |
| Outcome: | The Swiss Web Corpus for Applied Linguistics (SWS) is a multilingual collection of texts from selected web sources. |
VoxpopuliTTS: a large-scale multilingual TTS corpus for zero-shot speech generation (2025.coling-main)
Copied to clipboard
Wenrui Liu, Jionghao Bai, Xize Cheng, Jialong Zuo, Ziyue Jiang, Shengpeng Ji, Minghui Fang, Xiaoda Yang, Qian Yang, Zhou Zhao
| Challenge: | Existing multilingual TTS datasets are limited in speech generation fields due to lack of quality data. |
| Approach: | They propose to use 30,000 hours of high-quality speech data across 3 languages . they filter out low-quality text-text pairs and concatenate short transcripts . |
| Outcome: | The proposed dataset comprises 30,000 hours of high-quality speech data, across 3 languages with multiple speakers and styles, suitable for various speech tasks such as TTS and ASR. |